Back

Expert Systems with Applications

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Expert Systems with Applications's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Autonomous Spatial Transcriptomics Analysis (ASTA): Demonstrating Performance Improvements through Clustering, Biological Annotation, and AI-Driven Discovery

Zhang, M.; Roe, M.; Pollett, C.; Andreopoulos, W. B.

2026-08-18 bioinformatics 10.64898/2026.08.10.743848 medRxiv
Top 0.3%
1.0%
Show abstract

Spatial transcriptomics keeps measurement of gene expression while preserving spatial context, yet traditional analysis methods face challenges in computational efficiency, biological interpretability, and autonomous discovery. This project presents a framework solving these issues through three parts: (1) an ensemble clustering system achieving 66.7% improvement over baseline average and 23.9% over best single method with silhouette score of 0.540 and statistical significance (p = 0.0032, Cohens d = 1.82); (2) a knowledge-based clustering framework that annotates 88.6% of cells across 8 ovarian cell types using 428 marker genes; and (3) a GPT-4o-mini-powered autonomous agent that generated 3 biological hypotheses with validations.

2
Autonomous Loop Construction And Supervision For Clinician-Oriented Medical-Ai Research

Chen, X.; Jiang, X.; Shan, C.; Wang, Z.; Li, D.; Zhao, C.

2026-08-25 radiology and imaging 10.64898/2026.08.21.26361049 medRxiv
Top 0.3%
0.9%
Show abstract

Medical AI models have made a great impact on biomedical research and real-world clinical applications, but conducting interdisciplinary medical AI research remains challenging, requiring close collaboration between clinicians and AI experts. Recent advances in large language models (LLMs) and autonomous code agents present an opportunity for low cost medical AI development, where clinicians can build AI tailored to their own research questions, even without continuous support from dedicated AI experts. However, enabling code agents to autonomously tackle complex multimodal medical AI development tasks requires clinicians to construct and supervise an AI research loop with detailed technical specifics, demanding substantial expertise in AI and computer science that they often lack. To address this challenge, we introduce the Medical AI Research Loop Agent (MARLA), an agentic framework that completely abstracts the construction and supervision of medical AI research loops from clinicians. Given a clinician-defined research intent, MARLA automatically translates high-level research goals into executable hierarchical research loops, decomposes them into verifiable sub-loops, and specifies the models, datasets, tools, and evaluation protocols required for each task. During execution, MARLA coordinates specialized code agents, monitors progress, diagnoses failures, and iteratively refines research strategies based on experimental feedback to drive the research process toward optimal outcomes. We evaluate MARLA on multimodal medical AI tasks that require closed-loop conversion from high-level clinical study objectives to trained and validated AI models. The results demonstrate MARLA's ability to autonomously conduct complex medical AI research while substantially reducing the need for AI expertise.

3
CViT-ESP: Lightweight Pre-trained Vision Transformers for EEG-based Epileptic Seizure Prediction

Mohammad, U.; Parani, P.; Saeed, F.

2026-08-26 neuroscience 10.64898/2026.08.21.746341 medRxiv
Top 0.5%
0.6%
Show abstract

Background and Objective Epileptic seizure prediction is a critical challenge requiring the discrimination of subtle preictal physiological changes from interictal brain activity. While deep learning has shown promise in this domain, existing models often face limitations due to small EEG datasets, high computational costs for training from scratch, and a lack of patient-independent generalizability. In this paper, we present a novel framework for EEG-based seizure prediction that leverages pre-trained Vision Transformers (ViTs) through custom architectural modifications and optimized re-training strategies. Methods Our primary contributions include: [bullet]CVIT-ESP: A family of vision transformer architectures that replaces standard patch embedding layers with custom N-dimensional CNN stages to refine EEG representations. [bullet] ESPFormer: A lightweight, custom-designed transformer specifically engineered to mitigate overfitting on limited-scale EEG datasets. We identified optimal fine-tuning combinations for transformer blocks by devising a heuristic search-space reduction strategy, significantly reducing the training complexity. We validated our methods using the patient-independent MLSPred-Bench, involving 12 diverse benchmarks with varying seizure prediction horizons. Results Results demonstrate a clear progression in performance: while prior ResNet and vanilla Transformer models achieved an AUC-ROC of 69.0%, our CVIT-ESP architectures achieved the highest performance with a maximum average AUC of 76.4%. Conclusions These findings suggest that adapting pre-trained ViTs with domain-specific CNN front-ends and strategic fine-tuning offers a robust, generalizable, and resource-efficient path forward for clinical seizure prediction systems. Our code is available at: https://github.com/pcdslab/CVitEsp and https://github.com/pcdslab/ESPFormer

4
SAMP V2: A novel stacking ensemble learning model for antimicrobial peptides identification based on augmented split amino acid composition with biochemical-sequence-order information

Sun, M.; Wang, J.; Wan, S.

2026-08-19 bioinformatics 10.64898/2026.08.12.744552 medRxiv
Top 0.5%
0.6%
Show abstract

Antimicrobial resistance reduces the effectiveness of conventional antibiotics and has become a major global health threat, highlighting the need for new anti-infective agents. Antimicrobial peptides (AMPs), a diverse class of innate immune effectors with broad-spectrum antimicrobial activity, are promising candidates for combating drug-resistant infections. Identifying AMPs by wet-lab experiments, however, remains costly and time-consuming, creating a strong demand for computational identification methods. Our recently developed method, SAMP, captures region-specific residue distributions based on proportionalized split amino acid composition. However, SAMP might ignore key biochemical information and sequence order information. Here we present SAMP V2, a stacking ensemble learning framework based on biochemical and sequence-order information augmented split amino acid composition (BIA-SAAC), which extends SAMP by integrating pseudo-amino acid composition features with biochemical and sequence-order information into split peptide regions. Specifically, each peptide is divided into N-terminal, middle, and C-terminal regions, and pseudo amino acid composition is calculated within each region. Benchmarking tests on six independent test datasets, SAMP V2 outperformed multiple state-of-the-art models, including AMPpred-MFA and iAMP-Attenpred, in terms of accuracy, MCC, G-measure and F1-score. Given its high and robust performance, SAMP V2 could significantly accelerate the discovery of next-generation antimicrobial therapeutics for addressing the global threat of multidrug-resistant pathogens.

5
Clinically Generalisable End-to-End Graph Learning for CT Image-Based Multitask Stroke Diagnosis

Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26360026 medRxiv
Top 0.5%
0.6%
Show abstract

Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.

6
Predicting Protein-RNA Binding Affinity Changes via Spatial Coupling-Aware State Space Modeling

Chen, R.; Huang, X.; Jiang, H.; Ma, W.; Bi, X.; Wei, Z.; Nie, J.; Zhang, S.

2026-08-24 bioinformatics 10.64898/2026.08.23.745486 medRxiv
Top 0.6%
0.5%
Show abstract

Accurately predicting the effects of mutations on protein-RNA binding is crucial for elucidating disease mechanisms. Yet, exhaustively exploring the space of all possible variants is prohibitively expensive, motivating computational methods that can quantify mutation-induced changes in binding affinity (aka {Delta}{Delta}G) accurately and efficiently. We present iSCALE, an interpretable and generalizable deep learning method that adopts an implicit Spatial Coupling-Aware Ligand Encoding strategy to predict mutation-induced binding affinity changes. By injecting this implicit multiscale encoding scheme into a bidirectional state space modeling architecture, iSCALE learns a generalizable multiscale coupling pattern that achieves superior performances on not only the protein-RNA binding {Delta}{Delta}G, but also the protein stability {Delta}{Delta}G and protein-protein binding {Delta}{Delta}G predictions. Detailed analyses demonstrate that the model attention scores align well with structural characteristics. In addition, iSCALE shows good discriminative ability when predicting close samples such as complexes of same mutation but with different ligands or the same complex but with different mutation sites. In summary, iSCALE serves as an effective in silico tool for large-scale protein-RNA binding {Delta}{Delta}G prediction, which pushes the border of understanding in mutation-induced pathological outcomes.

7
Performance of a Self-Supervised Pretrained Neural Network for Orthopedic Radiograph Classification

Bagchi, R.; Yee, N. J.; Kwon, J. Y.; Taseh, A.; Ashkani-Esfahani, S.

2026-08-10 radiology and imaging 10.64898/2026.08.07.26359986 medRxiv
Top 0.6%
0.5%
Show abstract

Purpose To evaluate whether domain-adaptive self-supervised pretraining on musculoskeletal radiographs improves fracture classification and attribution faithfulness relative to ImageNet-pretrained baselines. Materials and Methods This study (June 2025 to May 2026) used previously acquired radiographs to compare three ResNet-50 initializations: supervised ImageNet pretraining (control), self-supervised ImageNet pretraining (DINO), and DINO with additional domain-adapted pretraining on 44,029 musculoskeletal radiographs (DINO-Ortho). All models underwent supervised fine-tuning in three experiments: in-distribution (MURA and FracAtlas datasets), out-of-distribution (an external dataset of 5,365 calcaneal radiographs from 1,775 patients), and initial weights (calcaneal radiographs only). Metrics included sensitivity, specificity, test accuracy, area under the receiver operating characteristic curve (AUROC), and Cohen's kappa; attribution faithfulness was quantified using Remove and Debias scores from Grad-CAM saliency maps. Comparisons used DeLong and Friedman tests. Results Classification performance did not differ significantly between DINO-Ortho and either baseline in any experiment (DINO-Ortho AUROC, 0.89 in-distribution and 0.95 with initial weights). All three models discriminated poorly out-of-distribution (control, 0.59; DINO, 0.57; DINO-Ortho, 0.58). DINO-Ortho showed significantly higher attribution faithfulness than both baselines in all three experiments, including out-of-distribution (25.39 vs -10.41 and 2.14; P < .001) and initial weights (20.88 vs 11.51 and 1.27; P < .001). Qualitative rankings favored DINO-Ortho but did not differ significantly. Conclusion Domain-adapted self-supervised pretraining on musculoskeletal radiographs improved attribution faithfulness while maintaining classification performance comparable to ImageNet-pretrained baselines; no model generalized adequately to external radiographs without task-specific fine-tuning.

8
Improving the Performance of Models Trained on Small EHR-Derived Samples by Leveraging External Data with Continual Learning Methods

Hui, J.; Xia, M.; Wilson, J.; Hill, E. D.; Scheer, A.; Franz, L.; Engelhard, M. M.; Goldstein, B. A.

2026-08-10 health informatics 10.64898/2026.08.08.26360010 medRxiv
Top 0.7%
0.5%
Show abstract

The performance of an EHR-based deep learning model trained on a small sample can be improved if more data is collected. Instead of collecting more data, the model can be trained on additional data from an analogous external source. However, this risks the model learning patterns in the external data that do not generalize to the target sample. Furthermore, data use agreements often prohibit combining datasets with medical records of different sources. We consider utilizing pre-existing methods in continual learning, namely the elastic weight consolidation (EWC) loss function and variational continual learning (VCL), both of which are regularization-based methods that we use to borrow external data and incorporate parameters from a model on external data into local model training. To investigate the utility of this modeling framework, we consider two binary classification tasks: (1) predicting which children will be diagnosed with autism spectrum disorder (ASD) from medical claims up to 18 months, and (2) predicting which patients with end-stage renal disease (ESRD) will be re-hospitalized within 30 days. Target datasets were derived from Duke University's EHR warehouse, and external datasets were sourced from either NC Medicaid claims for the ASD prediction task, or the United States Renal Data System (USRDS) for the rehospitalization prediction task. For both of these tasks, borrowing models - using either the EWC loss function or VCL - performed similarly to that of a model trained only on the full external data, when the sample size of target data used to train the model was small. That is, while a model that does not borrow using our methods performed poorly in low data regimes, the borrowing model instead matched the performance of a model trained on external data even when sample size of target data was small. In addition, an analysis of model predictions showed that models with small samples are better calibrated and more functionally similar to a model trained only on external data when the sample size is small.

9
ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation

Moran, J.; Freda, P. J.; Ghosh, A.; Hernandez, M. E.; Moore, J. H.

2026-08-25 bioinformatics 10.64898/2026.08.20.746074 medRxiv
Top 0.7%
0.4%
Show abstract

Summary: Agentic AI platforms enable the engineering of autonomous workflows but are not designed for experimentation and hypothesis testing. ASAREE (Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation), is an open-source platform to address this gap. ASAREE creates agents, connects to MCP servers and tools, and designs factorial experiments through a visual interface or Python SDK. It records a full provenance trace for every run and routes all model calls through a provider-agnostic bridge that supports local deployments, ensuring data privacy. As a use-case, we use ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline. Across a 2 x 2 x 2 factorial design, more advanced models, greater reasoning effort, and critic agent use significantly increased compute time, token use, cost, and feature count without improving predictive performance. The lowest-cost baseline, Claude Sonnet 5 with medium effort and no critic, achieved the highest mean PR AUC while Claude Opus 5 with extra high effort and a critic agent cost 15.5x more (USD) and ran 13.1x longer while performing worse on average. These findings highlight ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.

10
Report-Guided Semi-Supervised Learning for Scalable Prostate Cancer Detection on Biparametric MRI: Multicenter Prospective Validation and Multimodal Integration

Calado, A.; de Almeida, J. G.; Verde, A. S. C.; Tsiknakis, M.; Marias, K.; Regge, D.; Papanikolaou, N.; ProCAncer-I Consortium,

2026-08-07 radiology and imaging 10.64898/2026.08.05.26359781 medRxiv
Top 0.8%
0.4%
Show abstract

Purpose: To prospectively validate a semi-supervised learning framework with a lesion-only teacher model (RG-SSL-LOC) for scalable clinically significant prostate cancer detection on biparametric MRI (bpMRI) and assess its added value in multimodal models. Materials and Methods: A multicenter dataset of 13,706 bpMRI examinations (13,630 patients, 27 centers) was used for model development/validation. Three segmentation models (fully supervised learning [FSL], a state-of-the-art report-guided semi-supervised approach [RG-SSL], and the proposed RG-SSL-LOC) were evaluated at lesion- and case-level on external retrospective, external prospective, and internal prospective cohorts. Predictions from the best-performing model were combined with clinico-radiologic variables in a multimodal approach. All case-level results were compared with PI-RADS. Results: At lesion level, RG-SSL-LOC achieved higher median Dice than FSL and RG-SSL (0.49 vs 0.41 and 0.40; both p<.001). At case level, RG-SSL-LOC achieved area-under-the-curve (AUC) values of 0.83, 0.82, and 0.87 in the external retrospective, external prospective, and internal prospective cohorts, respectively. Compared with FSL, AUCs were 0.84 (p=.237), 0.80 (p=.020), and 0.84 (p<.001); compared with RG-SSL, AUCs were 0.83 (p=.929), 0.82 (p=.652), and 0.86 (p=.007); compared with PI-RADS, AUCs were 0.78 (p=.055), 0.83 (p=.652) and 0.86 (p=.480). Combined with clinico-radiological variables, RG-SSL-LOC significantly improved AUC versus clinico-radiological variables alone in the external retrospective (0.85 vs 0.80, p=.002), external prospective (0.87 vs 0.84, p=.008), and internal prospective (0.91 vs 0.88, p<.001) cohorts; in the latter, it reduced unnecessary biopsies by 15.19%. Conclusion: RG-SSL-LOC achieves better segmentation quality than other methods, demonstrates robust prospective multicenter performance and improves multimodal detection.

11
Quantifying User Engagement with the Helpilepsy Seizure Diary

Davies, J.; Biondi, A.; Viana, P. F.; Ampe, L.; Schreiber, J.; Richardson, M. P.

2026-08-07 health informatics 10.64898/2026.08.05.26359796 medRxiv
Top 0.8%
0.4%
Show abstract

Seizure diaries are one of the most useful sources of information in the management of epilepsy, however patient engagement with them can be sporadic. Sustained participation with seizure diaries affects the completeness and reliability of self-reported data, so it is vital to be able to measure engagement. To facilitate this, we create a multidimensional engagement metric with which to characterize how patients interact with their seizure diary. We utilise data from the Helpilepsy, a seizure diary application, common features found in application engagement metrics in business settings, and well understood clinical features to do this. Clustering is then performed to isolate different user groups based on how engaged they are, and these groups are studied to understand what drives the differences in engagement. We found three groups emerge from the clustering: low, medium and highly engaged users. Investigating these groups further, we put together a ``profile" for highly-engaged users. We find that they tend to be older at the point of diagnosis, and have had epilepsy for longer than the other users. We also find they tend to have had more medications, have higher doses of common anti-seizure medications, and they have more medications typically given to those with refractory epilepsy. The implications for e-diary design are that more attention should be given to those newer to epilepsy in the onboarding phase. Also, engagement is not necessarily based on just the upload of seizures, with other features of an e-diary being important to be filled in.

12
Does Data Preprocessing Affect Tree-Based Super Learners? An Investigation of Ensemble Optimization and Oracle Properties in Clinical Classification.

Darko, R.; Dwumah, D.; Agyapong, K. S.; Agyenim-Boateng, Y.; Darko Anim, R.; Wisdom Jakper, J.; Owusu-Ansah, N. K.; Owusu-Ansah, R.

2026-08-24 health informatics 10.64898/2026.08.20.26360880 medRxiv
Top 0.9%
0.3%
Show abstract

Machine learning workflows frequently incorporate data preprocessing to enhance predictive performance. However, the need for Super Learner ensembles made up only of preprocessing-invariant tree-based algorithms remains unexplored. Using three benchmark clinical classification datasets, this study examined how preprocessing affected the Super Learner's prediction performance, learner weight distribution, and oracle behavior. The Heart Disease (207 observations), Indian Liver Patient Dataset (583 observations), and Pima Indians Diabetes (768 observations) datasets were used to create a Super Learner ensemble model that included Classification and Regression Trees (CART), Random Forest, Ranger, and Extreme Gradient Boosting (XGBoost). Models were evaluated under raw and preprocessed data conditions using repeated cross-validation. Predictive performance was assessed using the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), and Brier score. Learner weight allocation and Oracle Gap were compared using paired Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Preprocessing produced negligible changes in predictive performance for the Heart Disease and Pima datasets. For the ILPD dataset, preprocessing significantly improved AUC (0.746 to 0.752; adjusted p = 0.0017) and reduced the Brier score (0.177 to 0.175; adjusted p < 0.001). Learner weights remained largely stable, although Random Forest replaced Ranger as the dominant learner for the Heart Disease dataset. Oracle Gaps remained extremely small (<0.002) across all datasets and did not differ significantly between preprocessing conditions. Preprocessing provides limited benefit for Super Learner ensembles composed of preprocessing-invariant learners and does not materially alter their oracle behavior. Preprocessing decisions should therefore be guided by dataset characteristics rather than adopted as a universal modelling practice.

13
PerturbTrace: Evaluating Feedback Use by AI Co-Scientist Agents in Perturbation Discovery

Yu, C.; Liu, S.; Qiao, G.; Luo, M.; Xiang, Y.; Xu, Z.

2026-08-20 bioinformatics 10.64898/2026.08.18.745260 medRxiv
Top 0.9%
0.3%
Show abstract

Recent advances in AI co-scientists have brought LLM agents into closed-loop experimental design. However, whether these agents use feedback from earlier rounds to revise subsequent experimental decisions remains unclear. We address this question with PerturbTrace, which evaluates each round-to-round transition through Feedback-to-State, State-to-Action, and Action-to-Outcome. These stages assess whether feedback is reflected in the agent's rationale and perturbation-selection strategy, whether the stated strategy guides the next perturbation batch, and whether that batch yields more hits than expected under random sampling. We evaluate four LLM agents on 17 screen-derived tasks and compare them with random selection, active learning, and LLM-guided Bayesian optimization baselines. Each agent outperforms the strongest non-agent method on at least 15 of the 17 tasks, yet controlled evaluations across six tasks show no consistent advantage from true feedback over random or no feedback. Among 576 transitions under true or random feedback, only 43 (7.5%) complete the full Feedback-State-Action-Outcome sequence, including 25 under random feedback. These findings show that high final recall does not necessarily indicate effective feedback use. They also highlight the need to evaluate closed-loop scientific agents by both their discovery performance and whether feedback changes their subsequent decisions.

14
SEEG Contact Detector: A 3D Slicer Extension for Automated Localisation of Intracranial Electrode Contacts

Smid, J.; Jezdik, P.; Kalina, A.; Kudr, M.; Janca, R.

2026-08-17 radiology and imaging 10.64898/2026.08.13.26360270 medRxiv
Top 1.0%
0.3%
Show abstract

Background: Precise localisation of intracranial electrode contacts is essential for the interpretation of stereoelectroencephalography recordings and planning epilepsy surgery. In current clinical practice, this is typically a manual process, which is time-consuming and prone to variability. Existing automated solutions are often fragmented across multiple tools requiring technical expertise, limiting their adoption in routine clinical workflows. This study presents an open-source extension for 3D Slicer that provides an integrated, user-friendly standalone solution for the direct automatic detection of electrode contacts within a widely used medical imaging platform. Results: The proposed method combines anchor bolt-based initialisation, probabilistic segmentation of electrode structures, and non-linear modelling to precisely track true electrode trajectories. The approach was evaluated on a dataset comprising 78 cases from 73 patients, including 1,078 electrodes with 14,480 contacts. The method achieved high localisation accuracy, with a median (interquartile range) deviation of 0.10 (0.06, 0.15) mm. Only 7/1078 (0.65%) electrodes required manual correction; these specific cases were handled using tools provided within the proposed extension. Conclusions: The presented extension enables fast, accurate, and reproducible electrode contact localisation within a single integrated environment. By combining automation with intuitive user interaction, it significantly reduces processing time while maintaining clinical reliability. The tool's free availability as an extension in 3D Slicer lowers the barrier to adoption and supports the standardisation of workflows across clinical and research centres.

15
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 1.0%
0.3%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

16
Offline Reinforcement Learning for Out-of-Distribution ICU Sepsis Decision Support

Arasteh, E.; Mirian, M. S.; Tavakol, M.

2026-08-25 health informatics 10.64898/2026.08.22.26361090 medRxiv
Top 1%
0.3%
Show abstract

Offline reinforcement learning (RL) provides a promising framework for learning and evaluating treatment policies from logged clinical data, particularly in sequential decision-making settings where prospective exploration would be unsafe. In ICU sepsis management, however, it remains unclear whether offline RL policies retain stable behavior under increasingly severe out-of-distribution (OOD) patient cohorts. In this paper, we evaluate standard offline RL methods on three severity-enriched OOD test mixtures from the MIMIC-III benchmark dataset to determine whether offline policies retain a stable, actionsensitive decision-support signal. Under the shared learned-dynamics offpolicy evaluation (OPE) protocol, as the severe-OOD ratio increases from 25% to 75%, observed clinical survival declines from 67% to 49%, while the best offline method in each mixture receives model-predicted terminal survival values of 87%, 86%, and 85%, respectively. Because observed clinical survival and model-predicted terminal survival are different quantities, this contrast suggests a stable model-based decision-support signal under severity shift. We further present a secondary physiological stabilization analysis using an episode-level physiological stabilization score (EPSS), a heuristic summary of whether selected physiological variables move in favorable directions during follow-up. In this analysis, model-generated rollouts under offline policies receive higher EPSS values than matched logged clinical trajectories for several physiological components. Together, these results support learned-dynamics OPE as a useful severity-OOD stress test for offline RL policies in ICU sepsis, while leaving prospective and causal validation as necessary next steps.

17
Multimodal artificial intelligence for personalized hepatocellular carcinoma treatment strategy selection

Feng, W.; Liu, S.; Yang, Z.; Tao, Y.; Gu, X.; Jin, W.

2026-08-25 health informatics 10.64898/2026.08.21.26361067 medRxiv
Top 1%
0.3%
Show abstract

Background Hepatocellular carcinoma (HCC) treatment selection demands nuanced integration of heterogeneous patient data, yet prevailing predictive models rely on restricted data modalities and oversimplified therapeutic frameworks, compromising clinical translation. Objective We developed and validated a multimodal artificial intelligence framework to guide optimal treatment strategy selection across the full spectrum of HCC interventions. Methods This retrospective study comprised 1,043 HCC patients (development cohort, January 2017-December 2023) and 55 external validation patients (2023) from Wuxi Peoples Hospital. We engineered Embedding-Augmented Extra Trees (ET-Emb), a novel model fusing structured clinical variables with contextual text embeddings derived from medical histories and radiology reports. ET-Emb quantifies probabilities for five primary treatments: open/laparoscopic resection, transarterial chemoembolization, radiofrequency ablation (RFA), and chemotherapy. Model performance was rigorously assessed via 10-fold cross-validation and external validation using ROC-AUC and PR-AUC metrics. Results ET-Emb demonstrated robust performance in the development cohort (ROC-AUC: 0.84 {+/-} 0.04; PR-AUC: 0.55 {+/-} 0.06), significantly outperforming established benchmarks. This generalizability was preserved in external validation (ROC-AUC: 0.77 {+/-} 0.02; PR-AUC: 0.47 {+/-} 0.03). SHAP analysis identified textual clinical narratives and socioeconomic determinants as critical predictive drivers. Conclusions By unifying structured and unstructured data modalities, ET-Emb delivers accurate, multi-treatment strategy prediction for HCC. Its clinical validity and the demonstrated significance of textual features establish multimodal AI as an essential paradigm for simulating complex oncological decision-making, positioning ET-Emb as a transformative tool for precision HCC management.

18
A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Abhigyan, R.; Sood, V.; Arora, P.; Kaur, B.

2026-08-09 bioinformatics 10.64898/2026.08.04.742799 medRxiv
Top 1%
0.2%
Show abstract

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative-evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. HighlightsO_LIProposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. C_LIO_LIPerformed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. C_LIO_LICombined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. C_LIO_LIDemonstrated the applicability of the proposed framework across multiple bioactive peptide datasets. C_LI

19
MaternaAI: Enhancing Equitable Maternal Healthcare in Kerala with Fairness-Aware and Explainable Learning Models

Jo, A. A.

2026-08-14 obstetrics and gynecology 10.64898/2026.08.12.26360340 medRxiv
Top 1%
0.2%
Show abstract

Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.

20
Towards Interpretable AI Second Opinions: Foundation Model Heatmaps in Radiology

Dack, E.; Dai, C.; Hoppe, H.; Krueselmann, P.; Meiler, S.; Jutidamrongphan, W.; Wang, L.; Tang, K.

2026-08-23 radiology and imaging 10.64898/2026.08.20.26360908 medRxiv
Top 1%
0.2%
Show abstract

AI-assisted diagnostic tools typically act as a "second opinion," providing radiologists with a discrete prediction or probability score that can be consulted alongside clinical context. This treats AI as an independent advisor rather than a collaborative partner, leaving its reasoning largely opaque. We explore a complementary approach grounded in human-AI collaboration through visual interpretability. Specifically, we investigate (1) radiologist performance when diagnosing chest X-rays from images alone, and (2) whether deep learning-generated heatmaps can support radiologists during this diagnostic process, rather than merely validating a final answer. We developed an interactive application that enables readers to engage directly with model-generated heatmaps as they form their diagnoses, and conducted a user study to evaluate how this influences diagnostic behaviour and accuracy. Our findings offer new insights into integrating interpretable, spatially grounded AI feedback into radiologist workflows. Code, datasets, and the application can be found at https://github.com/eedack01/heatmap_assisted_diagnosis.